Anthropic2026-08-25 11:41:10Claude test leaks point to strong 3D reasoning in two unreleased modelsTwo unreleased Claude models, referred to in public testing discussions as "Marshmallow" and "Melon," are drawing attention after early hands-on results surfaced online. According to the material cited in the source report, the models showed unusually strong performance in 3D reinforcement learning, architectural spatial planning, and the generation of complex spatial relationships, with some outputs reportedly completed in a single pass rather than through repeated prompt revisions. The report says developers spotted a model ID, claude-marshmallow-ht-eap, being called 57 times in API traffic between 00:45 and 00:57Z on Aug. 21. It later appeared in Claude Code as a "Custom model" with a 1 million-token context window. At the time of publication, neither model appeared to have a first-party API endpoint, suggesting access may have been limited to red-team testers or internal users. Another detail repeated by testers was the heavy use of "thinking tokens." One X user said both models consumed so many of these internal reasoning tokens that tests repeatedly hit the <max-tokens> limit. The timing has added to speculation around Anthropic’s roadmap, coming less than a month after the July 24 release of Opus 5 and following Sonnet 5 on June 30.1330